Skip to main content

System Monitoring

Overview

The system monitoring screen shows whether the Security365 service is currently operating normally.OverviewThis is a screen that can be done.

You can check the status of server resources (CPU·Memory·Disk·Network), the operating status of each service, and the status of shared infrastructure (Elasticsearch·RabbitMQ·Redis·Storage) on a single page.

  • All information on this screen isRead-onlyis.
  • data isAutomatically refresh every 60 secondsIt is possible.
  • This feature is**Provided only in an on-premises (customer dedicated installation) environment.**It is possible.

When data is not visible, the data on this screen is provided by the internal data collection server of the system. If the collection server is not ready or communication is interrupted, the screen may be empty or display "No Data." If this state persists, please contact the person in charge.

View at a glance

The screen is composed of a single page, arranged in the following order from top to bottom.

areaContent
Top BarAuto Update Status, Last Update Time, Refresh Button, Export Button
Summary CardSummarize the status of systems, services, and resources in numbers
Infrastructure HealthElasticsearch·RabbitMQ·Redis·Storage Status Card
Tab AreaService Status(Default Screen) andSystem Monitoring(Resource Trends) Two Tabs

Top Bar

At the top of the page, there are buttons for data refresh and export.

itemMeaning · Usage
Automatic UpdateIf it is turned on, all data will be automatically refreshed every 60 seconds. If turned off, you need to press the refresh button to update.
Last updatedDisplays the time when the data was last updated.
Refresh ButtonFetching the latest data immediately. (Cannot click again for a moment.)
Export ButtonDownload the data on the current screen as an Excel file. (See "Export Data" below)

If data updates are delayed or interrupted, in the automatic update place**"Data Update Delay"or"Data Update Suspension"**A guide will be displayed and the screen will be blurred. It will automatically refresh when the connection is restored.

Summary Card

The summary card at the top of the screen quickly shows the overall operational status in numbers. Hovering over the information (i) icon next to each card title will display a description.

System Status

It shows the number of server (system) hardware statuses based on CPU, memory, and disk usage rates.

displaymeaning
Normal N개Number of servers operating within the normal range
Warning N itemsNumber of servers that require attention due to high usage
障害 N개Number of servers with very high usage or issues

If the number of warnings and failures is greater than 0, you can check what server it is in the "System Monitoring" tab below.

Service Status

It shows the number of currently running security services that are operating normally.

displaymeaning
Normal N개Number of services running normally
Warning N itemsNumber of services with some anomalies detected but operation is ongoing
障害 N개Number of services that are interrupted or have recurring errors

If there is a warning or outage, check which service it is in the "Service Status" tab.

Maximum Resource Usage Status

Shows the highest resource usage figures among all servers. It is used to quickly identify the most heavily loaded points.

itemmeaning
Network ReceptionMaximum Network Ingress (Mbps) Across All Servers
CPUMaximum CPU Usage (%) Among All Servers + Status Badge
MemoryMaximum Memory Usage (%) of All Servers + Status Badge

Meaning of Status Colors (Common)

The gauge and badge colors across the screen follow the criteria below.

ColorUsage Rate Criteriameaning
Normal (Blue/Green)Less than 70%Relaxed steady state
Warning (Orange)70% or more ~ less than 90%High usage state that requires attention
Risk (Red)over 90%Load is very high. Response delays and other impacts may occur, so inspection is recommended.

Infrastructure Health

Displays the status of the shared infrastructure used by security services in a card format.**If one fails, multiple products can be affected simultaneously.**These are important services. Clicking each card will open a detailed popup.

cardroleInformation displayed on the card
ElasticsearchData RepositorySystem count (Normal/Total)
RabbitMQMessage BrokerNumber of consumers, Number of unprocessed messages (cases)
RedisCache and Session StorageRole (Master/Replica), Memory Usage
StorageStorage SpaceUsage Rate (%) Gauge, Usage/Total Capacity (GB)

Each card displays the status badge on the right as Normal, Warning, or Error, and if data cannot be retrieved,**"No Data"**It will be displayed in (gray).

If the storage usage exceeds 90%, there may be issues with logging and data recording, so please contact the person in charge if it is marked in red.

Service Status Tab (Default Screen)

Displays in a table format which server each service is running on and its status.

Filter

You can filter to see only the desired items using the three filters at the top.

FilterAction
ServiceSearch and Select by Service Name
systemSelect a specific server (node) to display only the services of that server.
statusCheck and display only the desired status among Normal, Warning, and Error.

If there are no items that meet the criteria**"There is no service data to display."**It will be displayed.

Two Service Groups

  • Product Service: SHIELD DRM, SHIELD Gate, etc. Security365 core security services
  • Data Service: Shared infrastructure used by multiple products (Elasticsearch, RabbitMQ, Redis)

How to View Service Cards

At the top of each service group card, a summary is displayed:Instance N(total execution unit count),Normal N / Warning N / Failure N. If there is a problem, a warning icon will be displayed along with a reason such as "Restart N times".Detailed InformationPressing the button will open the detailed popup.

The table below is structured as follows.

rowmeaning
Composition ServiceIndividual components (rows) that make up the service
(Server Name) ColumnStatus badge on each server (node). If not running on that server, "-", unassigned server is "Unassigned"
Maximum Resource UtilizationMaximum CPU and memory usage of the service configuration (e.g., "CPU 0.12 core · MEM 512 MB")
Reason for abnormalityDisplay the cause of the problem in text when there is an issue.

Meaning of abnormal reasons

displaymeaningRecommended Actions
Service InterruptionThe configuration service is not operational.Immediate contact with the person in charge
Mobilization StandbyStarting the serviceCheck again shortly. If it lasts long, please contact us.
Restart Repeat (N times)Error restarting repeatedly (based on 24 hours). Highlight if more than 3 times.Possibility of recurrence. Contact person inquiry
Response DelaySlow service responseLoad and Network Inspection

Detailed Information Popup

Product Service Details— Opens when the service is clicked.

  • Displays the number of instances and the summary of normal/warning/failure at the top.
  • There are cards that expand by component (workload), and items with issues automatically expand.
  • per instance (execution unit)CPU·Memory·Uptime·Restart Count·Statusshows it in a table.
  • An instruction message is displayed along with instances that have issues.
Guide textmeaning
Waiting for activation.Still starting out
The process has been terminated.Execution stopped
The process is restarting due to repeated crashes. (Restart N times)Restarting due to error — inspection needed
The container is not ready yet.Started but service preparation is not complete.

Data Service Details— Elasticsearch / RabbitMQ / Redis card opens when clicked.

  • The key metrics for each service are displayed along with their status in the top summary.
ServiceIndicators displayed in detail
ElasticsearchSystem count (Normal/Total)
RabbitMQNumber of consumers, Number of unprocessed messages
RedisRole (Master/Replica), Memory Usage
  • Below, the instance-specific status (instance, system, CPU, memory, uptime, restart, status) is displayed in a table.

You can check the time flow of resource usage by server in a chart, and the current usage in a card.

Time Range

buttonQuery Period
1 hourLast 1 hour
6 hoursRecent 6 hours (default)
24 hoursRecent 24 hours
Recent 7 daysRecent 7 days

The 24-hour interval is summarized at approximately 5-minute intervals, and the recent 7-day interval is summarized at approximately 30-minute intervals. This is normal operation.

How to Read Charts

CPU Usage (%), Memory Usage (%), Disk Usage (%), Network In (Mbps) four charts are displayed.

  • Each server is indicated by a line of a different color.
  • When you hover over the chart, the exact figures for each server at that time will be displayed.
  • Clicking on the server name in the legend below will hide or show the individual server lines.
  • If the usage rate surged at a specific time, check the service status at the same time as well.

While loading data, "Loading data...", if there is no data, "No data", and if the query fails, "Unable to view time series data." will be displayed.

System Usage Rate Card

Below the chart, the current usage status by server is displayed in cards, and clicking on a card willNode Detail PopupThis opens.

itemDisplay content
Server Name · IPServer Identification Information
Operating TimeElapsed time since last restart (e.g., "12 days 5 hours")
NetworkTransmission·Reception (Mbps)
CPU / Memory / DiskUsage Rate (%) Gauge + Details (CPU is used/total cores, Memory·Disk is used/total GB)

If the uptime is very short (within a few hours), the server was recently restarted.

Node Detail Popupwith the above resource statusList of services running on the serverYou can check (service·severity·number of instances·number of restarts). If the number of restarts is 3 or more, a warning icon will be displayed along with it.

Status Badge Overview

The status badges throughout the screen mean the following.

badgemeaning
● NormalNormal operation
▲ WarningSome anomalies detected, but operation is ongoing (high usage or restarted more than 3 times)
■ ErrorStopped or repeated errors occurred — inspection needed
No data (gray)Unable to check the status (did not receive data from the collection server)

Connection Status and Data Refresh

  • If automatic updates are enabled, all data will be refreshed every 60 seconds.
  • If the data refresh is delayed, "Data Update Delayed" will be displayed, and if it is interrupted, "Data Update Interrupted" will be shown, and the screen will be dimmed. The displayed values may not be the latest.
  • When the connection is restored, it automatically returns to the "Normal Refresh" state, and the data is updated immediately.
  • If the "Data Update Suspension" continues, please contact the person in charge.

Data Export

topExportPressing the button will download the current screen's data as an Excel file.

itemContent
File FormatExcel file(.xlsx) 1 piece (2 sheets)
Sheet 1 — System MonitoringResources by Server (Hostname·IP·Operating System·Uptime·CPU·Memory·Disk·Network Transmission and Reception)
Sheet 2 — Service Statusservice·namespace·node·status·container readiness·CPU·memory·restart count·uptime
file namesecurity365_monitoring_{날짜_시각}.xlsx

When the export is complete, the message "Download is complete." will be displayed. (The PDF operational report is in preparation.)

Frequently Seen Messages (Guidance by Problem Situation)

MessageMeaning / Action
Loading data...Loading data. Please wait a moment.
No data / No data to display.No data to display. If the issue persists, check the status of the collection server.
There is no service data to display.There are no services that match the criteria. Please check the filters.
Data Update Delay / InterruptionData update is delayed or halted. If it continues, please contact the person in charge.
Cannot check the service status.Temporary error. Please try again later.
Cannot check system resources.Temporary error. Please try again later.
Failed to create file.Export failed. Please try again later.